Papers with neural-based metrics
Uncertainty-Aware Machine Translation Evaluation (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Several neural-based metrics have been proposed to evaluate machine translation quality, but they are trained on noisy, biased and scarce human judgements. |
| Approach: | They propose a method to evaluate machine translation quality using point estimates . they combine COMET framework with Monte Carlo dropout and deep ensembles . |
| Outcome: | The proposed methods perform well across multiple language pairs and with references. |
Towards Multiple References Era – Addressing Data Leakage and Limited Reference Diversity in Machine Translation Evaluation (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent research shows a weak correlation between n-gram-based metrics and human evaluations in machine translation tasks. |
| Approach: | They propose to use multiple references generated by LLMs to improve alignment between automatic metrics and human evaluations. |
| Outcome: | The proposed approach improves the alignment between automatic metrics and human evaluations on the WMT22 benchmark with 4 languages and achieves a maximum accuracy gain of 9.5%. |
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization (2025.acl-long)
Copied to clipboard
| Challenge: | n-gram-based metrics are considered indicative (even if imperfect) of human evaluation for English, but their suitability for other languages remains unclear. |
| Approach: | They systematically assess evaluation metrics for generation for languages and tasks using n-gram-based and neural-based metrics. |
| Outcome: | The proposed evaluation suite is based on eight languages from four typological families and shows that it is sensitivity to the language type at hand. |